An Evolution Strategies Approach to the Simultaneous Discretization of Numeric Attributes
نویسندگان
چکیده
Many data mining and machine learning algorithms require databases in which objects are described by discrete attributes. However, it is very common that the attributes are in the ratio or interval scales. In order to apply these algorithms, the original attributes must be transformed into the nominal or ordinal scale via discretization. An appropriate transformation is crucial because of the large influence on the results obtained from data mining procedures. This paper presents a hybrid technique for the simultaneous supervised discretization of continuous attributes, based on Evolutionary Algorithms, in particular, Evolution Strategies (ES), which is combined with Rough Set Theory and Information Theory. The purpose is to construct a discretization scheme for all continuous attributes simultaneously (i.e. global) in such a way that class predictability is maximized w.r.t the discrete classes generated for the predictor variables. The ES approach is applied to 17 public data sets and the results are compared with classical discretization methods. ES-based discretization not only outperforms these methods, but leads to much simpler data models and is able to discover irrelevant attributes. These features are not present in classical discretization techniques.
منابع مشابه
Chi2: feature selection and discretization of numeric attributes
Discretization can turn numeric attributes into discrete ones. Feature selection can eliminate some irrelevant attributes. This paper describes Chi2, a simple and general algorithm that uses the 2 statistic to discretize numeric attributes repeatedly until some inconsistencies are found in the data, and achieves feature selection via discretization. The empirical results demonstrate that Chi2 i...
متن کاملMining Frequent Ranges of Numeric Attributes via Ant Colony Optimization for Continuous Domains without Specifying Minimum Support
Currently, all search algorithms which use discretization of numeric attributes for numeric association rule mining, work in the way that the original distribution of the numeric attributes will be lost. This issue leads to loss of information, so that the association rules which are generated through this process are not precise and accurate. Based on this fact, algorithms which can natively h...
متن کاملAn Iterative Improvement Approach for the Discretization of Numeric Attributes in Bayesian Classifiers
The Bayesian classifier is a simple approach to classification that produces results that are easy for people to interpret. In many cases, the Bayesian classifier is at least as accurate as much more sophisticated learning algorithms that produce results that are more difficult for people to interpret. To use numeric attributes with Bayesian classifier often requires the attribute values to be ...
متن کاملFeature Selection via Discretization
| Discretization can turn numeric attributes into discrete ones. Feature selection can eliminate some irrelevant and/or redundant attributes. Chi2 is a simple and general algorithm that uses the 2 statistic to discretize numeric attributes repeatedly until some inconsistencies are found in the data. It achieves feature selection via dis-cretization. It can handle mixed attributes, work with mul...
متن کاملDiscretizing Continuous Attributes Using Information Theory
Many classification algorithms require that training examples contain only discrete values. In order to use these algorithms when some attributes have continuous numeric values, the numeric attributes must be converted into discrete ones. This paper describes a new way of discretizing numeric values using information theory. The amount of information each interval gives to the target attribute ...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2003